Why Did DeepSeek's Long Text Suddenly Act Weird? ByteDance Seed Team Unveils the Mystery of AI Performance Fluctuations
In a recent paper, the Seed team at ByteDance revealed that performance fluctuations in large language models when processing ultra-long texts are mainly caused by the phase sensitivity of block-based KV cache compression technology. To reduce memory consumption during long-context reasoning, this technique compresses continuous token windows into fewer entries by using a fixed stride, but it introduces new position coordinates for tokens relative to the compressed window, causing retrieval bias and leading to performance fluctuations.